Papers with Social Media
SEDTWik: Segmentation-based Event Detection from Tweets Using Wikipedia (N19-3)
Copied to clipboard
| Challenge: | Recent work on event detection from tweets has focused on localized events or breaking news only. |
| Approach: | They propose to split tweets into segments, extract bursty segments, cluster them, summarize them. |
| Outcome: | The proposed system can detect newsworthy events occurring at different locations of the world from a wide range of categories. |
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (2026.findings-eacl)
Copied to clipboard
| Challenge: | Social media platforms such as X (formerly Twitter), Facebook, and Reddit generate user-generated content. |
| Approach: | They propose a framework to assess privacy risks in social media by evaluating vulnerabilities across six dimensions: data collection, preprocessing, visibility, fairness, computational risk, and regulatory compliance. |
| Outcome: | The proposed framework assesses privacy risks across six dimensions . it achieves F1-scores of 0.58–0.84, but incurs 1% - 23% drop under fine-tuning . |
Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection (2021.acl-short)
Copied to clipboard
| Challenge: | a lack of labeled, non-English resources for hate speech detection limits research on hate speech . a recent study shows that zero-shot, cross-lingual learning models cannot be used as they are . lack of consistency limits research, and lack of models for non-english languages limits learning . |
| Approach: | They propose a zero-shot, cross-lingual transfer learning framework for hate speech detection . they use benchmark data sets in English, Italian, and Spanish to detect hate speech . |
| Outcome: | The proposed framework can't be used as it is, but needs to be carefully designed, the authors say . they find that non-hateful, language-specific taboo interjections are misinterpreted as signals of hate speech . |
MediaHG: Rethinking Eye-catchy Features in Social Media Headline Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Creating a good headline on social media platforms requires a disentanglement-based model to balance the content and contextual features. |
| Approach: | They propose a disentanglement-based headline generation model which can balance the content and contextual features by incorporating contrastive learning and auxiliary multi-tasking to choose the best domain-suitable headline. |
| Outcome: | The proposed model can balance content and contextual features, while allowing bloggers to obtain more site traffic and profits while readers can have easier access to topics of interest. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)
Copied to clipboard
Irina Temnikova, Iva Marinova, Silvia Gargova, Ruslana Margova, Alexander Komarov, Tsvetelina Stefanova, Veneta Kireva, Dimana Vyatrova, Nevena Grigorova, Yordan Mandevski, Stefan Minkov
| Challenge: | SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis. |
| Approach: | They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them. |
| Outcome: | The proposed package is the first to be released in Bulgarian and is available for free. |